Papers with test-time search

2 papers
AgentRM: Enhancing Agent Generalization with Reward Modeling (2025.acl-long)

Copied to clipboard

Challenge: Existing LLM-based agents have strong performance on held-in tasks, but their generalizability to unseen tasks remains poor.
Approach: They propose a reward-based generalizable reward model to guide the policy model for effective test-time search.
Outcome: The proposed agentRM outperforms existing agents on held-in tasks by 8.8 points on average.
CoTEvol: Self-Evolving Chain-of-Thoughts for Data Synthesis in Mathematical Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to improve LLMs' reasoning abilities suffer from diminishing returns or high computing overhead.
Approach: They propose a genetic evolutionary framework that casts CoT generation as a population-based search over reasoning trajectories.
Outcome: The proposed framework improves correct-CoT synthesis success by over 30% and enhances structural diversity with markedly improved efficiency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations